The problem with accuracy
Accuracy is defined as the fraction of correct predictions out of all predictions. Simple, intuitive, easy to explain. It is also deeply misleading in many of the situations where you most need a reliable metric.
Consider a fraud detection system. In a typical dataset, maybe 0.5% of transactions are fraudulent. A model that predicts "not fraud" for every single transaction achieves 99.5% accuracy. It catches zero fraudulent transactions. Nobody would use it. Yet accuracy says it is excellent.
Imagine rating a smoke alarm purely by how often it is silent. A smoke alarm that never beeps achieves a perfect silence rate. But it also fails to do its one job: alerting you when there is a fire. Accuracy on imbalanced datasets is the smoke alarm version of model evaluation. It measures the wrong thing entirely.
The moment your classes are imbalanced, which they almost always are in real fraud, medical, and security applications, you need better tools. This lesson gives you those tools.
Classification metrics that actually matter
Everything you need starts from the confusion matrix. You met it in Lesson 3.2. Now let's build the key metrics from its four cells: True Positives (TP), True Negatives (TN), False Positives (FP), and False Negatives (FN).
The precision-recall tradeoff
Precision and recall are locked in a tug-of-war. If you make your model more aggressive about flagging positives, recall goes up (you catch more real cases) but precision goes down (you also raise more false alarms). If you make it more conservative, precision goes up but recall falls.
Which end of that tradeoff you prefer depends entirely on the cost of each type of error in your specific context.
| Context | Worse error | Optimise for | Why |
|---|---|---|---|
| Cancer screening | False Negative (missed cancer) | High Recall | Missing a real cancer case has far greater consequences than unnecessary follow-up tests. |
| Spam filtering | False Positive (real email flagged as spam) | High Precision | Users accept some spam in their inbox more readily than they accept missing important emails. |
| Fraud detection | False Negative (missed fraud) | High Recall | Each missed fraud costs money. False positives mean some legitimate transactions get reviewed manually. |
| Content recommendation | False Positive (irrelevant recommendation) | High Precision | Showing bad recommendations erodes user trust quickly. It is better to show fewer but more relevant items. |
When you cannot clearly prioritise one over the other, use F1. When your problem has a heavy class imbalance, consider also reporting the balanced accuracy or the Matthews Correlation Coefficient, which account for the imbalance directly.
The ROC curve and AUC-ROC
Most classifiers do not just output a class label. They output a probability. "This email has a 91% chance of being spam." You then apply a threshold: anything above 0.5 gets labelled spam. But that threshold is a choice. Different thresholds give different precision-recall tradeoffs.
The ROC (Receiver Operating Characteristic) curve visualises the full range of this tradeoff across every possible threshold. It plots True Positive Rate (recall) against False Positive Rate (1 minus specificity) as the threshold sweeps from 0 to 1. A perfect classifier hugs the top-left corner. A random classifier follows the diagonal.
The further the curve bends toward the top-left corner, the better the classifier at separating classes regardless of threshold. AUC (Area Under the Curve) summarises this into a single number you can use to compare models directly.
AUC-ROC is especially useful because it is threshold-independent. You can compare two models on their AUC without committing to a particular decision threshold upfront. Once you have selected the best model by AUC, you then choose the threshold based on your precision-recall priorities for the deployment context.
AUC = 1.0: Perfect. The model correctly separates every positive from every negative. AUC = 0.9: Excellent. The model would correctly rank a random positive above a random negative 90% of the time. AUC = 0.7: Acceptable for some tasks; investigate further. AUC = 0.5: Random guessing. Your model has learned nothing useful.
Regression evaluation metrics
Regression problems have their own set of metrics, because error in continuous outputs cannot be measured as right or wrong. Every prediction has a magnitude of error, and different metrics weight those magnitudes differently.
When you report regression results to a non-technical audience, MAE is your friend. "Our model predicts delivery time with an average error of 8 minutes" is something any stakeholder can understand immediately. RMSE is harder to interpret but better for comparing models with each other because it penalises bad predictions heavily. R-squared gives the most complete picture but can be misleading on non-linear data.
Putting it all together in code
Here is a single Python snippet that generates a full evaluation report for a classification model: accuracy, precision, recall, F1, and the ROC-AUC score. Copy this pattern and use it on every classification project.
from sklearn.datasets import load_breast_cancer from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split from sklearn.metrics import ( accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, classification_report ) # Load data and train a model X, y = load_breast_cancer(return_X_y=True) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y ) model = RandomForestClassifier(n_estimators=100, random_state=42) model.fit(X_train, y_train) # Hard predictions (class labels) preds = model.predict(X_test) # Soft predictions (probabilities) for AUC probs = model.predict_proba(X_test)[:, 1] # Print full report print("=== Classification Report ===") print(f"Accuracy: {accuracy_score(y_test, preds):.4f}") print(f"Precision: {precision_score(y_test, preds):.4f}") print(f"Recall: {recall_score(y_test, preds):.4f}") print(f"F1 Score: {f1_score(y_test, preds):.4f}") print(f"AUC-ROC: {roc_auc_score(y_test, probs):.4f}") print() print(classification_report(y_test, preds, target_names=['Malignant', 'Benign']))
Accuracy: 0.9649
Precision: 0.9730
Recall: 0.9730
F1 Score: 0.9730
AUC-ROC: 0.9955
precision recall f1-score support
Malignant 0.95 0.95 0.95 42
Benign 0.97 0.97 0.97 72
The AUC-ROC of 0.9955 is exceptional. Even if we lower the decision threshold to catch more malignant cases at the cost of some false alarms, this model will almost certainly outperform any reasonable baseline. The F1 scores for both classes are above 0.95, confirming the model is not just winning because one class dominates the data.
You have finished Phase 3: Machine Learning
You started this phase not knowing what a model actually does. You now understand the learning loop, classification, regression, clustering, how overfitting works and how to fight it, and how to measure whether a model is genuinely useful. That is a solid foundation. Phase 4 takes everything you have built and extends it into neural networks and deep learning.